Summary:
Conventionally, synthetic training data quality is evaluated through human perception, prioritizing visual realism. From the model’s perspective, what truly matters is whether a sample lies within the right region of its embedding space. This work introduces VERSE, a methodology for analyzing and improving the performance of Vision–Language Models by exploring their visual embedding space. VERSE enables the visualization of latent representations to assess model feasibility, identifies problematic regions, and guides synthetic data generation to enhance performance in those clusters. We validate the proposed methodology for Visually-rich Document Understanding by training on the synthetic MERIT Dataset and evaluating on its real-world counterpart, MERIT Secret, focusing on key information extraction as a sequence-generation task scoped to transcripts of records in Spanish. Results show that VERSE uncovers the visual features associated with error-prone clusters, and that retraining with samples containing these features substantially boosts F1 performance without degrading generalization. On-premise models optimized with VERSE—Donut (F1 = 0.76) and Idefics2 (F1 = 0.81)—match or surpass SaaS solutions such as GPT-4o (F1 = 0.78) and Pixtral (F1 = 0.73), preserving data privacy and avoiding external APIs.
Spanish layman's summary:
VERSE es una metodología para analizar y mejorar VLMs explorando su espacio de embeddings visuales. Identifica clústeres propensos a error y guía la generación de datos sintéticos, logrando que modelos on-premise (Donut, Idefics2) igualen o superen a soluciones SaaS como GPT-4o preservando la privacidad.
English layman's summary:
VERSE is a methodology for analyzing and improving VLMs by exploring their visual embedding space. It identifies error-prone clusters and guides synthetic data generation, enabling on-premise models (Donut, Idefics2) to match or surpass SaaS solutions like GPT-4o while preserving data privacy.
Keywords: Visually-rich Document Understanding; Vision-Language Models; Visual embeddings; Interpretability; Explainability
JCR-JIF Impact Factor and WoS quartile: 9,100 - Q1 (2025)
DOI reference:
https://doi.org/10.1016/j.patcog.2026.114448
Published on paper: December 2026.
Published on-line: July 2026.
Citation:
I. de Rodrigo, A.J. López López, J. Boal, "VERSE: Visual Embedding Reduction and Space Exploration - Latent-space clustering for improving document understanding", Pattern Recognition, Vol. 180, nº. Part D, pp. 114448, December 2026. [Online: July 2026] doi: 10.1016/j.patcog.2026.114448